Tag
2 articles
OpenAI claims GPT-5.6 Sol outperforms Opus 5 on ARC-AGI-3, but only in its own custom test environment. The official benchmark shows a stark contrast in performance.
Anthropic's Opus 5, combined with Auto Mode, shows zero success rate in preventing browser-based prompt injection attacks, a major AI security vulnerability.